Back

Sleep Advances

Oxford University Press (OUP)

All preprints, ranked by how well they match Sleep Advances's content profile, based on 11 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
Does the Sleep Regularity Questionnaire capture objective sleep-wake regularity? Evidence from wearable and sleep diary data.

Driller, M. W.; Bodner, M. E.; Fenuta, A.; Stevenson, S.; Suppiah, H.

2026-02-26 health informatics 10.64898/2026.02.24.26347047 medRxiv
Top 0.1%
56.0%
Show abstract

Sleep regularity is an important but under-measured dimension of sleep health. Objective indices from actigraphy or wearables are robust but resource-intensive. The Sleep Regularity Questionnaire (SRQ) offers a brief subjective tool, but its validity against objective and diary-based indices in healthy adults is unclear. In Part 1, 31 adults wore a smart ring continuously for 21 nights. Device-derived regularity metrics included the Sleep Regularity Index (SRI), interdaily stability (IS), social jetlag (SJL), composite phase deviation (CPD), and the standard deviation of sleep onset and wake time. In Part 2, 52 adults completed a one-week sleep diary, from which variability in sleep timing, total sleep time (TST), SJL and nightly perceived sleep quality were derived. All participants completed the SRQ and Brief Pittsburgh Sleep Quality Index (B-PSQI). In Part 1, associations between SRQ scores and device-derived SRI, IS, SJL, CPD and timing variability were small (absolute r [≤] 0.36). Higher SRQ Global and Sleep Continuity scores were moderately associated with better B-PSQI global scores (r -0.37 to -0.44). In Part 2, SRQ Global and Circadian Regularity showed small-to-moderate associations with higher diary-rated sleep quality and lower bedtime variability (r {approx} 0.40 and -0.32 to -0.34), while correlations with other diary metrics and B-PSQI were weak (absolute r [≤] 0.25). The SRQ shows modest convergent validity with diary-based timing variability and perceived sleep quality, but only weak correspondence with smart ring-based sleep regularity indices. It is likely to complement, rather than replace, objective monitoring in healthy adults with relatively regular sleep-wake patterns.

2
Multivariate determinants of wearable-measured sleep quality across a large observational cohort: roles of physical activity, gut microbiome, blood analytes, and lifestyle factors.

Cavon, J.; Perez, C.; Quinn-Bohmann, N.; Magis, A. T.; Gibbons, S. M.

2026-05-29 health informatics 10.64898/2026.05.27.26354250 medRxiv
Top 0.1%
49.0%
Show abstract

Emerging evidence links the gut microbiome to sleep quality, yet measuring sleep at scale remains challenging. Commercial wearables, such as Fitbit, capture objective sleep and activity data in naturalistic settings. We integrated Fitbit data from a large, deeply-phenotyped cohort with paired lifestyle and health questionnaires. Wearable-derived measures aligned well with self-reported sleep, activity, and happiness. We identified dozens of covariate-adjusted associations between Fitbit-derived sleep features, lifestyle factors, and multi-omic data. Among molecular feature sets, the gut microbiome showed the greatest number of associations with sleep quality: butyrate-producing genera were positively associated with sleep and amplified the benefits of physical activity. Oscillospira, in particular, was consistently associated with better sleep. In blood, insulin, omega-3, and cortisol correlated with poorer sleep, whereas lower alcohol intake and mineral supplements correlated with better sleep. These robust, covariate-adjusted findings advance mechanistic understanding of the gut-sleep axis and broader molecular and lifestyle determinants of sleep quality.

3
Naturalistic sleep tracking in a longitudinal cohort: how long is long enough?

Goparaju, B.; De Palma, G.; Bianchi, M. T.

2024-10-21 health informatics 10.1101/2024.10.19.24315818 medRxiv
Top 0.1%
45.6%
Show abstract

BackgroundDespite broad interest in the health implications of sleep duration, traditional measurements via polysomnography or actigraphy are often limited to one or a few nights per person. Given the potential variability of sleep duration over time, inferential uncertainty remains an important issue for relatively short observation windows. MethodsWe describe potential limitations of shorter duration sleep tracking by sub-sampling from longer-term observation windows, using a combined approach of simulated data from known distributions, in addition to real-world data (30-365 nights) from over 35,000 participants who provided informed consent to participate in the Apple Heart and Movement Study and elected to contribute sleep data to the study. ResultsSimulations demonstrate that the magnitude of deviation from truth, defined using all available observations per individual, as well as the presence and direction of bias, depended on the sub-sample size, the type of simulated distribution (Gaussian versus skewed), and the summary statistics of interest, such as central tendency (mean, median) and dispersion (standard deviation (SD), interquartile range). For example, the SD computed from n=7 observations from a simulated normal distribution (7+1 hours) showed a median 6.7% under-estimation bias (IQR 24% under- to 14.7% over-estimation). Real-world sleep duration data, when under-sampled and compared to longer observations within-participant, showed similar SD bias at 7 nights, and similar convergence rates approaching the true value (based on 90 nights) as longitdunal sample number increases. Shapiro-Wilk tests for normality and log-normality show that 64% of simulated log-normal (skew) distributions fail to reject normality at n=7 samples, while real-world sleep duration data most commonly failed both normality and log-normality tests. Finally, simulated cohorts with sleep durations of 7+1 hours mixed with a subset of 6+1 hours sleepers showed that a random single-night observation of "short sleep" (6 hours) is more likely from random variation of a 7-hour sleeper, than from an actual 6-hour sleeper. Extending the observation to n=7 nights mitigates this mis-classification risk. ConclusionThe results of simulations and empiric data patterns suggests that longer duration tracking provides important and tangible benefits to reduce bias and uncertainty in sleep health research that historically relies on small observation windows.

4
The effect of physical activity timing on insomnia and sleep quality: a randomized cross-over trial in older adults

Albalak, G.; Noordam, R.; van der Elst, M.; Drop, T.; Caneda Cabrera, E.; Oudendijk, L.; Lammers, G. J.; Gordijn, M.; Kervezee, L.; Exadaktylos, V.; van Bodegom, D.; van Heemst, D.

2026-05-20 geriatric medicine 10.64898/2026.05.18.26353463 medRxiv
Top 0.1%
44.0%
Show abstract

Background Insomnia symptoms are common in older adults. While observational studies suggest physical activity (PA) timing affects health outcomes, its effect on sleep remains unclear. We compared morning versus evening PA effects on insomnia severity and sleep quality in older adults with insomnia symptoms. Methods Eligible participants were aged 60 to 80 years with (sub)clinical insomnia (Insomnia Severity Index [ISI] score [≥]10). In a randomized cross-over trial, participants engaged in coached PA in the morning (10:00 - 11:00) or evening (19:30 - 20:30) for 14 days each. ISI scores were assessed post-intervention. Objective sleep parameters; duration, latency, efficiency, and timing, were assessed with a Withings Sleep Analyzer under the mattress. Subjective sleep quality was reported daily via smartphone app. Salivary dim light melatonin onset (DLMO) was measured on the final day of each intervention. Results Of 37 participants (mean ISI 14.3 {+/-} 3.3), 27 completed the study (mean age 69.8 {+/-} 5; 63% women). ISI scores improved after both morning ({Delta} - 2.5; 95% CI: - 1.14, - 3.83) and evening ({Delta} - 2.0; 95% CI: - 0.63, - 3.38) activity relative to baseline, but were not different between interventions. Compared to evening activity, sleep midpoint occurred earlier with morning activity (03:40 vs 04:00; {Delta} - 20 min; 95% CI: - 31, - 8). No differences in subjective sleep quality or DLMO were found. Exploratory analyses suggested insomnia scores improved specifically in late chronotypes following morning activity. Conclusions While morning vs. evening PA timing did not impact most sleep quality measures, it influenced sleep timing. Larger studies are needed to define optimal and personalized PA timing for improving sleep.

5
Promoting Sleep Duration in the Pediatric Setting Using a Mobile Health Platform: A Randomized Optimization Trial

Mitchell, J. A.; Morales, K.; Williamson, A.; Jawahar, A.; Juste, L.; Vajravelu, M. E.; Zemel, B.; Dinges, D.; Fiks, A.

2023-01-05 pediatrics 10.1101/2023.01.04.23284151 medRxiv
Top 0.1%
42.6%
Show abstract

ObjectiveDetermine the optimal combination of digital health intervention component settings that increase average sleep duration by [&ge;]30 minutes per weeknight. MethodsOptimization trial using a 25 factorial design. The trial included 2 week run-in, 7 week intervention, and 2 week follow-up periods. Typically developing children aged 9-12y, with weeknight sleep duration <8.5 hours were enrolled (N=97). All received sleep monitoring and performance feedback. The five candidate intervention components (with their settings to which participants were randomized) were: 1) sleep goal (guideline-based or personalized); 2) screen time reduction messaging (inactive or active); 3) daily routine establishing messaging (inactive or active); 4) child-directed loss-framed financial incentive (inactive or active); and 5) caregiver-directed loss-framed financial incentive (inactive or active). The primary outcome was weeknight sleep duration (hours per night). The optimization criterion was: [&ge;]30 minutes average increase in sleep duration on weeknights. ResultsAverage baseline sleep duration was 7.7 hours per night. The highest ranked combination included the core intervention plus the following intervention components: sleep goal (either setting was effective), caregiver-directed loss-framed incentive, messaging to reduce screen time, and messaging to establish daily routines. This combination increased weeknight sleep duration by an average of 39.6 (95% CI: 36.0, 43.1) minutes during the intervention period and by 33.2 (95% CI: 28.9, 37.4) minutes during the follow-up period. ConclusionsOptimal combinations of digital health intervention component settings were identified that effectively increased weeknight sleep duration. This could be a valuable remote patient monitoring approach to treat insufficient sleep in the pediatric setting.

6
Hidden in the Night: Wearable Sleep Assessment of Nocturnal Hypoglycaemia in Type 1 Diabetes

Alsuhaymi, A.; Nutter, P. W.; Thabit, H.; Harper, S.

2026-01-28 health informatics 10.64898/2026.01.22.26344161 medRxiv
Top 0.1%
40.4%
Show abstract

BackgroundNocturnal hypoglycaemia (NH) is a common and challenging complication in Type 1 Diabetes (T1D), disrupting blood glucose control and sleep physiology. Its real-world impact on sleep architecture remains poorly characterised. Consumer wearables offer a way to examine these associations under free-living conditions, providing detailed insight into behavioural and physiological responses to nocturnal blood glucose fluctuations. This study aims to assess how wearable-derived sleep metrics and physiological features could be used as indicators of NH, including the effects of how low blood glucose levels fall during hypoglycaemic events and the associated pre-event changes. MethodsWe conducted a comparative observational analysis of paired continuous glucose monitoring (CGM) and Garmin smartwatch data collected over 12 weeks from 17 adults with T1D. Nights were categorised as normoglycaemia, hyperglycaemia, or hypoglycaemia Level 1 ([&ge;]3.1 and <3.9 mmol/L), and hypoglycaemia Level 2 (<3.0 mmol/L). Thirteen sleep metrics, including total sleep time, wake after sleep onset (WASO), sleep-stage proportions, fragmentation indices, and physiological features such as heart rate, were compared using non-parametric tests. Pre-hypoglycaemic event analyses examined 60-minute and 15-minute windows preceding hypoglycaemia to identify early deviations in sleep and physiological metrics. ResultsAcross 573 nights, 17.5% involved Level 1 and 7.3% Level 2 hypoglycaemia. Level 2 hypoglycaemia was associated with 31 minutes less wakefulness, 17-25 minutes more REM, and up to 74% more deep sleep compared with normo-glycaemic nights. Sleep efficiency increased during hypoglycaemic events despite greater fragmentation. Pre-hypoglycaemic episode analyses revealed shorter awake and light-sleep bouts, as well as a 9.8% higher heart rate, preceding Level 2 episodes. ConclusionsWearable-derived sleep and physiological signals reveal clear intraindividual changes both before and during NH. Our findings indicate that Level 2 episodes are associated with deeper sleep and reduced behavioural arousal, suggesting that CGM alarms may be less effective at waking individuals during level2 NH. By characterising pre-hypoglycaemic changes that differ based on hypoglycaemia level, this work provides preliminary evidence for personalised, wearable-based early-warning systems. Such approaches could help distinguish nocturnal hypoglycaemic events and support more effective alerting, particularly in settings with limited or no access to CGM. Author SummaryO_ST_ABSWhy was this study done?C_ST_ABSPeople with Type 1 Diabetes (T1D) frequently experience nocturnal hypoglycaemia (low blood glucose at night), a dangerous event that often goes unnoticed because individuals are less able to recognise symptoms or wake up during sleep. These events also disrupt sleep in ways that are not well characterised under real-world conditions. Limited access to continuous glucose monitoring (CGM), especially in low- and middle-income countries, highlights the need for affordable alternatives to ensure nighttime safety. What did we do and find?Using more than 500 nights of paired smartwatch and CGM data, we investigated how sleep features change when blood glucose levels fall overnight. We found that hypoglycaemic nights show distinct alterations in sleep architecture, including increased REM and deep sleep, and greater micro-fragmentation. A key finding was that Level 2 hypoglycaemia was associated with deeper sleep and reduced wakefulness. This pattern indicates that individuals may be less likely to awaken during more severe events, even when alarms are present. Pre-hypoglycaemic episode analysis revealed additional early-warning signals, such as shorter awake and light-sleep bouts and elevated heart rate, before level 2 hypoglycaemia occurred. What do these findings mean?Smartwatches can capture sleep-based changes that appear before and during nocturnal hypoglycaemia. Because deeper sleep during Level 2 episodes may reduce responsiveness to CGM alerts, these results suggest that current alarm approaches could be improved by incorporating sleep features alongside glucose data. Such sleep-informed detection may enhance the reliability of hypoglycaemia alerts, reduce missed events during deep sleep, and provide a foundation for low-cost early-warning systems in settings where CGM is unavailable or unaffordable. Further research is needed in larger and more diverse populations, but this work provides early evidence that wearable-derived sleep features can meaningfully strengthen nocturnal hypoglycaemia detection.

7
Severity of Depression and Anxiety Symptoms Manifest in Physiological and Behavioral Metrics Collected from a Consumer-Grade Wearable Ring

Sameh, A.; Azadifar, S.; Nauha, L.; Karmeniemi, M.; Niemela, M.; Farrahi, V.

2026-02-09 health informatics 10.64898/2026.02.06.26345566 medRxiv
Top 0.1%
40.2%
Show abstract

Wearable devices can collect changes in human behaviors related to mental health including depression and anxiety. Here, we examined whether and how digital metrics from a consumer-grade wearable smart ring (Oura Ring) differed by severity of depression and anxiety symptoms using data from a large-scale population-based sample of young adults (n=1,290, age range: 33-35). Participants wore the ring for two weeks, assessing sleep architecture, nocturnal heart rate (HR), heart rate variability (HRV), and movement intensity. Mental health symptoms were assessed using the Generalized Anxiety Disorder 7-item and Hopkins Symptom Checklist-25 scales. On average, participants with higher depression and/or anxiety symptoms had lower levels of rapid eye movement and had higher levels of deep and light sleep, elevated nocturnal HR, reduced HRV, and lower daytime movement compared to non-symptom individuals. Findings suggest that symptoms of depression and anxiety may manifest in physiological and behavioral metrics collected by consumer-grade wearable devices.

8
Making sleep behaviors interpretable: adapting the two-process model of sleep regulation to longitudinal Fitbit sleep and activity behaviors for health insights

Coleman, P.; Annis, J.; Master, H.; Gustavson, D. E.; Han, L.; Brittain, E.; Ruderfer, D. M.

2026-03-03 health informatics 10.64898/2026.03.01.26347356 medRxiv
Top 0.1%
38.8%
Show abstract

BackgroundAs sleep data from wearable devices are increasingly available in health research, there are new opportunities to understand sleep regulation behaviors as modifiable risk factors for disease. At such a large scale (tens of thousands of people over millions of day-level observations), prioritizing and interpreting sleep behaviors is challenging while maintaining biological relevance and modifiability. In this work, we aim to address this challenge by proposing a framework to interpret Fitbit data through a well-known neurobiological framing of sleep regulation, the two-process model. MethodsWe use data from the All of Us Research Program, a national biobank with passively collected Fitbit data for 32,292 people across 15,754,893 total days. We map Fitbit behaviors (b) to either circadian (C) or homeostatic (S) processes. Using iterative exploratory factor analysis to obtain weights, the Fitbit Cb and Sb are then weighted at the level of each day to create Cb and Sb scores. FindingsCb and Sb scores were found to align with expected real-world relationships with age, seasonality, shift work, and napping. Cb and Sb scores were interpreted with relation to depression, where it was found that Sb scores are highly associated with likelihood of diagnosis (OR = 1.5, p < 2e-16) while Cb and Sb scores are equally associated with severity (Sb score {beta} = 0.2, Cb score {beta} = 0.21, p < 2e-16). InterpretationCb and Sb scores support longitudinal interpretation (e.g., changes in Sb around treatment), aggregation (e.g., differences in Cb between two groups), and actionable modification (e.g., reduce naps to improve poor Sb). Overall, our behavior scores allow for interpretation of wearables sleep data and can be utilized across many disease contexts to better understand how sleep influences health. FundingThis work was supported by NIH training grant T32GM145734 and NIH R21HL172038.

9
Estimating the sleep period time window based on a hip-worn accelerometer collected in children and adults

Migueles, J. H.; van Hees, V. T.; Stein, M. J.; Leitzmann, M. F.; Baurecht, H.; Lendt, C.

2025-11-27 health informatics 10.1101/2025.11.25.25340956 medRxiv
Top 0.1%
34.6%
Show abstract

BackgroundAccurately detecting the Sleep Period Time (SPT) window in the daily life is essential for understanding habitual sleep and health. Although actigraphy devices (accelerometers) placement varies across studies, most SPT-detection algorithms are developed for wrist data. Open-source algorithms support reproducibility and transparency in estimating the SPT. AimsTo optimise and evaluate two open-source algorithms, HDCZA and HorAngle, for estimating the SPT window using hip-worn accelerometer data. MethodsA total of 109 children and 194 adults wore wrist and hip accelerometers for six nights and completed sleep diaries. An established algorithm combining wrist and diary data served as the reference. HDCZA and HorAngle parameters were optimised using Bayesian optimisation on 60% of the sample and evaluated in the remaining 40%. ResultsMean differences for sleep onset and wake-up were -3 and 4 minutes for HDCZA (limits of agreement [LoA]: -221,215 and -185,194; root-mean square error [RMSE]=111 and 97) and 0 and -4 minutes for HorAngle (LoA: -199,199 and -223,214; RMSE=111 and 112). For SPT duration, mean differences were 7 minutes (LoA: -252,266; RMSE=132) for HDCZA and -4 minutes (LoA: -254,246; RMSE=128). No significant differences in SPT duration were found (P=0.774; P=0.237). Both algorithms showed moderate agreement with the reference in ranking sleep duration ({kappa} {approx} 0.56-0.58). Differences were unrelated to age or sex but linked to non-wear time. ConclusionsBoth open-source algorithms demonstrated value for estimating the SPT window from hip data. While HDCZA requires no additional sensor-specific parameters, HorAngle depends on accurate axis identification. Statement of SignificanceAccurately estimating the sleep period time (SPT) window from hip-worn accelerometers is essential for studies assessing sleep in free-living conditions. However, most available algorithms were developed for wrist-worn data. This study optimised and validated two open-source algorithms, HDCZA and HorAngle, for hip-worn accelerometer data in children and adults. Both algorithms performed comparably to a wrist-based reference using sleep diaries, showing consistent agreement across age and sex. These methods enable researchers to estimate habitual sleep without additional sensors or diaries, improving reproducibility and scalability in observational research. The algorithms are openly implemented in the GGIR R package, offering accessible and standardised tools for analysing hip-based accelerometer data.

10
Performance evaluation of an under-mattress sleep sensor versus polysomnography in >400 nights with healthy and unhealthy sleep

Manners, J.; Kemps, E.; Lechat, B.; Catcheside, P.; Eckert, D.; Scott, H.

2024-09-11 health informatics 10.1101/2024.09.09.24312921 medRxiv
Top 0.1%
31.0%
Show abstract

Consumer sleep trackers can provide useful insight into sleep and sleep patterns. However, large scale performance evaluation studies against direct sleep measures are needed to comprehensively understand sleep tracker accuracy. This study evaluated performance of an under-mattress sensor to estimate sleep and wake versus polysomnography, during multiple in-laboratory protocols in a large sample including individuals with and without sleep disorders and during day versus night sleep opportunities. 183 participants (51% male, mean[SD] age=45[18] years) attended the sleep laboratory for a research study that included simultaneous polysomnography and under-mattress sensor (Withings Sleep Analyzer [WSA]) recordings. Epoch-by-epoch analyses with confusion matrices were used to determine accuracy, sensitivity, and specificity of the WSA versus polysomnography. Bland-Altman plots examined bias in sleep duration, efficiency, onset-latency, and wake after sleep onset. Overall WSA sleep-wake classification accuracy was 83%, sensitivity 95%, and specificity 37%. The WSA significantly overestimated total sleep time (48[81]minutes), Sleep efficiency (9[15]%), sleep onset latency (6[26]), and underestimated wake after sleep onset (54[78]), p<0.05. Accuracy and specificity were higher for night versus daytime sleep opportunities in healthy individuals (89% and 47% versus 82% and 26% respectively, p<0.05). Accuracy and sensitivity were also higher for healthy individuals (89% and 97%) versus those with sleep disorders (81% and 91%, p<0.05). WSA performance is comparable to other consumer sleep trackers, with high sensitivity but poor specificity compared to polysomnography. Poorer accuracy and specificity during daytime versus night-time sleep opportunities is likely due to increased wake time and reduced sleep efficiency. Contactless, under-mattress sleep sensors show promise for accurate sleep monitoring, noting the tendency to over-estimate sleep particularly where wake time is high.

11
Predicting daily sleep outcomes from continuous HRV in female chronic pelvic pain disorders

Clarke, R.; Shahnawaz, S.; Hirten, R.; Rodrigues, J.; Landell, K.; Danieletto, M.; Ona, G.; Ensari, I.

2026-07-17 health informatics 10.64898/2026.07.16.26357390 medRxiv
Top 0.1%
23.0%
Show abstract

Background: Female chronic pelvic pain disorders (CPPDs) are highly prevalent and frequently accompanied by sleep disturbance and autonomic nervous system (ANS) dysregulation. Heart rate variability (HRV), a non-invasive index of ANS function, may provide an objective, physiological correlate of sleep health and can be monitored using wearable devices, enabling a continuous, scalable approach. Objectives: This study examined whether wearable-derived daily HRV metrics are associated with self-reported sleep disturbance in women with CPPD(s) compared with healthy controls, using epoch-level data and generalized additive models. Methods: We conducted a retrospective observational study using up to 90 days of data from a mobile health research app. Participants were 128 women with CPPD(s) and 63 demographically matched healthy controls, who completed a daily PROMIS-based 3-item sleep disturbance questionnaire and wore Fitbit devices that provided 5-minute HRV epochs. Primary predictors were high frequency (HF) and low frequency (LF) power and root mean square of successive differences (RMSSD), with group (CPPD vs control), daily pain severity, and menstrual status as covariates. We fit separate generalized additive mixed models (GAMMs) for each HRV metric with a nonlinear smooth term and an HRV x Group interaction. Results: Higher HF and RMSSD were associated with lower sleep disturbance scores, and these associations were stronger in controls than in the CPPD group (HF x group B {approx} -1.59, p < 0.00010; RMSSD x group B {approx} -0.58, p < 0.0001). LF showed a more complex pattern but also differed by group (B {approx} -0.531, p < 0.0001). HRV smooth terms were highly nonlinear, and models explained ~8-9% of deviance in sleep disturbances. Pain severity and menstrual bleeding were strongly associated with worse sleep. Conclusion: These findings indicate small but consistent associations between wearable-derived HRV metrics and daily sleep disturbances in women with CPPD(s) and healthy controls, with weaker associations in CPPD(s). Integrating continuous HRV with symptom tracking could support low-burden and multimodal monitoring of sleep health in chronic pelvic pain, but prospective validation is needed before HRV can be used for diagnostic or treatment response decision making.

12
Leveraging NLP to Identify Domain-Specific Variables in Large-Scale Cohort Metadata: A Sleep Use Case

Draper, B.; Briggs, P.; Purcell, S. M.; Dijk, D.-J.; Bauermeister, S.; Bartsch, U.

2026-01-19 health informatics 10.64898/2026.01.18.26344317 medRxiv
Top 0.1%
22.8%
Show abstract

Public health policies increasingly rely on the use of complex and large datasets containing heterogeneous, multimodal data that require advanced analytical methods to extract meaningful insights and support evidence-based decision-making. Essential for the sharing and analysis of public health data is the description of the data ("data about data") or metadata. Indeed, a lack of metadata standards has been identified as a key technical barrier to public health data sharing 1. Metadata varies considerably between cohort studies. It often contains extremely heterogenous variable and data descriptions for similar or identical metrics. The assessment and processing of metadata is very labour intensive. It relies on researchers manually sifting through large numbers of variable descriptors and other study data documentation to extract variables of interest. Unless a validated quantitative tool is used, the variability in phrasing is surprisingly high. Variations in metadata can makes it hard to find and compare results across studies. Specifically, questions about wake-sleep behaviour like subjective sleep quality and duration are commonly employed in large cohort studies - but vary not only in the wording of the questions put to participants but also in the metadata that describes the questions and their responses. We developed a semi-automatic retrieval method (METAMATCH) for sleep-related variables from metadata obtained from multiple large cohort studies. Here we employ sleep as the health area of interest, but in principle this method can be applied to any area of research. We developed the retrieval method using metadata provided by Dementias Platform UK (DPUK, https://www.dementiasplatform.uk/) a Trusted Research Environment (TRE) that hosts over 50 cohort datasets of different sizes. From this we extracted a metadata corpus describing 86,682 variables across 17 cohort studies. To identify sleep-related variables, we curated a reference dictionary (the SLEEPTALKING corpus) combining terms from the Unified Medical Language System (UMLS) and expert knowledge. We then employed Term Frequency-Inverse Document Frequency (TF-IDF) vectorisation and cosine-similarity to rank variable-descriptions by comparing the metadata corpus to the SLEEPTALKING corpus. We identified 337 semantically unique sleep-related variables. We then performed semantic analyses to characterize these variables. We applied Bidirectional Encoder Representations from Transformers (BERT) embeddings and unsupervised cluster analysis. The results indicate that definitions tend to group by cohort and length of the variable description. We validated the cluster analysis using expert-based categorisation of variable descriptions. The most common categories identified were Sleep Difficulty, Sleep Duration, and Sleep Latency & Waking. Consequently, expert knowledge and consensus remains essential for accurately categorizing variables within a universal ontology across different studies. We conclude that combining NLP techniques with expert input offers an efficient approach to harvesting metadata across multiple cohort studies.

13
A comparison of sleep metrics from mid-thigh and low-back accelerometers to wrist based data using open-source algorithms

Passfield, G.; Mackay, L.; Crofts, C.; Schofield, G.

2024-11-11 health informatics 10.1101/2024.11.10.24317079 medRxiv
Top 0.1%
22.5%
Show abstract

IntroductionWearable accelerometers are a valuable tool for monitoring sleep, sedentary behaviour, and physical activity patterns within 24h time-use in free-living environments. While wrist-worn accelerometers are favoured for monitoring sleep, they do not accurately distinguish between sitting and lying positions (Narayanan et al., 2020). This study aims to determine whether back or thigh-mounted accelerometers yield sleep metrics comparable to wrist-worn devices using an open-source algorithm originally validated for the wrist. MethodsData from 20 healthy sleepers were collected using Axivity AX3 accelerometers. Participants wore accelerometers on their right thigh, low-back, and wrist for one night of sleep in their own bed. Sleep metrics were calculated using the van Hees algorithm through the GGIR package in R. The primary outcomes were: Total Sleep Time (TST), Wake After Sleep Onset (WASO), Awakenings (AWK), Sleep Efficiency (SE), Sleep Interval (SI) and Sleep Onset Timestamp (SOT). Within-subject ANOVA with Tukeys post hoc, Pearson correlation coefficients, Bland-Altman plots, and Cohens d were used to assess the comparability of sleep metrics between the body placements. ResultsData analysis included all 20 participants. Mid-thigh accelerometers demonstrated a strong linear relationship with wrist accelerometers across all metrics (r = 0.86-0.98). Bland-Altman plots demonstrated a narrow 95% confidence interval suggesting that wrist and mid-thigh metrics are in good agreement, except for AWK which is slightly underestimated by the mid-thigh device. Conversely, low-back accelerometers demonstrated moderate linear relationship with the wrist (r = 0.63-0.98) and the Bland-Altman results showed wide limits of agreement with significant overestimations of TST, SE, SI and underestimations of WASO, AWK, SOT. Cohens d demonstrated small differences between mid-thigh and wrist devices, except for AWK (d= 0.42). Low-back values for WASO, SE, and AWK showed moderate differences. ConclusionsThis analysis demonstrates that the mid-thigh accelerometer yields comparable sleep metrics to wrist-worn devices when processed with the van Hees algorithm.

14
Unveiling Sleep Dysregulation in Chronic Fatigue Syndrome with and without Fibromyalgia Through Bayesian Networks

Bechny, M.; Scurati, M.; van der Meer, J.; Faraci, F.; Natelson, B.; Kishi, A.

2025-02-10 health informatics 10.1101/2025.02.06.25321788 medRxiv
Top 0.1%
22.2%
Show abstract

Chronic Fatigue Syndrome (CFS) and Fibromyalgia (FM) often co-occur as medically unexplained conditions linked to disrupted physiological regulation, including altered sleep. Building on the work of Kishi et al. [7], who identified differences in sleep-stage transitions in CFS and CFS+FM females, we exploited the same strictly controlled clinical cohort using a Bayesian Network (BN) to quantify detailed patterns of sleep and its dynamics. Our BN confirmed that sleep transitions are best described as a second-order process [14], achieving a next-stage predictive accuracy of 70.6%, validated on two independent data sets with domain shifts (60.1-69.8% accuracy). Notably, we demonstrated that sleep dynamics can reveal the actual diagnoses. Our BN successfully differentiated healthy, CFS, and CFS+FM individuals, achieving an AUROC of 75.4%. Using interventions, we quantified sleep alterations attributable specifically to CFS and CFS+FM, identifying changes in stage prevalence, durations, and first- and second-order transitions. These findings reveal novel markers for CFS and CFS+FM in early-to-mid-adulthood females, offering insights into their physiological mechanisms and supporting their clinical differentiation.

15
Light exposure during sleep is associated with irregular sleep timing: the Multi-Ethnic Study of Atherosclerosis (MESA)

Wallace, D. A.; Qiu, X.; Schwartz, J.; Scheer, F. A.; Redline, S.; Sofer, T.

2023-10-12 occupational and environmental health 10.1101/2023.10.11.23296889 medRxiv
Top 0.1%
19.7%
Show abstract

ObjectiveExposure to light at night (LAN) may influence sleep timing and regularity. Here, we test whether greater light exposure during sleep (LEDS) associates with greater irregularity in sleep onset timing in a large cohort of older adults. MethodsLight exposure and activity patterns, measured via wrist-worn actigraphy (ActiWatch Spectrum), were analyzed in 1,933 participants with 6+ valid days of data in the Multi-Ethnic Study of Atherosclerosis (MESA) Exam 5 Sleep Study. Summary measures of LEDS averaged across nights were evaluated in linear and logistic regression analyses to test the association with standard deviation (SD) in sleep onset timing (continuous variable) and irregular sleep onset timing (SD[&ge;]1.36 hours, binary). Night-to-night associations between LEDS and absolute differences in nightly sleep onset timing were also evaluated with distributed lag non-linear models and mixed models. ResultsIn between-individual linear and logistic models adjusted for demographic, health, and seasonal factors, every 5-lux unit increase in LEDS was associated with an increase of 7.8 minutes in sleep onset SD ({beta}=0.13 hours, 95%CI:0.09-0.17) and 40% greater odds (OR=1.40, 95%CI:1.24-1.60) of irregular sleep onset. In within-individual night-to-night mixed model analyses, every 5-lux unit increase in LEDS the night prior (lag0) was associated with a 2.2-minute greater deviation of sleep onset the next night ({beta}=0.036 hours, p<0.05). Conversely, every 1-hour increase in sleep deviation (lag0) was associated with a 0.35-lux increase in future LEDS ({beta}=0.347 lux, p<0.05). ConclusionLEDS was associated with greater irregularity in sleep onset in between-individual analyses and subsequent deviation in sleep timing in within-individual analyses, supporting a role for LEDS in exacerbating irregular sleep onset timing. Greater deviation in sleep onset was also associated with greater future LEDS, suggesting a bidirectional relationship. Maintaining a dark sleeping environment and preventing LEDS may promote sleep regularity and following a regular sleep schedule may limit LEDS.

16
Electrodermal Activity as a Critical Modality for Wearable Sleep Monitoring: A Comprehensive Systematic Review from Fundamental Physiology to Clinical Translation

Suraparaju, P. V.; Udhayakumar, S.

2025-11-17 health informatics 10.1101/2025.11.14.25340258 medRxiv
Top 0.1%
19.7%
Show abstract

Wearable sleep monitoring devices have proliferated over the past decade, driven by consumer interest in sleep optimization and athletic recovery tracking. However, current consumer-grade wearables suffer from fundamental accuracy limitations, with meta-analysis of 798 patients across 24 studies showing wrist-worn devices systematically underestimate rapid eye movement (REM) sleep by 50-70%, with error rates exceeding 2 hours per night in some cases. Photoplethysmography (PPG)-based heart rate variability represents the dominant approach in current wearables, achieving only 60-72% accuracy for four-stage sleep classification. Electrodermal activity (EDA), a pure sympathetic nervous system marker, offers complementary physiological information previously unexploited in wearable devices. This comprehensive systematic review of 87 peer-reviewed studies involving 2,015 subjects across 1,847 separate sleep recordings synthesizes three critical findings: (1) Wrist EDA physiology during sleep fundamentally diverges from daytime conventions, exhibiting 86-91% nights of superior amplitude compared to palm measurements on 84-91% of nights, contrary to established anatomical hierarchy; (2) Wrist versus fingertip EDA measurement reveals opposing site-specific advantages during sleep, with wrist showing 2.02-2.35x higher amplitude, 34% fewer motion artifacts, 68% lower electrode drift variability, and 89.2x stronger sleep stage discrimination effect; (3) Multimodal integration of wrist EDA with PPG, accelerometry, and temperature increases four-stage sleep classification accuracy from 72% to 83% (11 percentage point improvement), while EDA-based machine learning achieves 83.7% accuracy for clinically relevant sleep apnea screening - a potential 2 billion dollar annual market opportunity. The wrist location provides practical manufacturing advantages (34% cost reduction for EDA subsystem, 7-6 dollars per unit savings) while fundamentally overturning decades of measurement conventions and establishing the physiological and practical basis for next-generation wearable sleep architecture. This analysis consolidates emerging evidence into an actionable roadmap for translating EDA into consumer and clinical wearable devices.

17
Performance of a Semi-Automated Hierarchical Rest Interval Detection Pipeline (actiSleep) for Wrist Actigraphy in Adolescents

Soehner, A. M.; Kissel, N.; Hasler, B. P.; Franzen, P. L.; Levenson, J. C.; Clark, D. B.; Buysse, D. J.; Wallace, M. L.

2026-03-06 psychiatry and clinical psychology 10.64898/2026.03.05.26347744 medRxiv
Top 0.1%
19.6%
Show abstract

Actigraphy is a popular behavioral sleep assessment tool in research and clinical practice. Hierarchical hand-scoring approaches remain the standard for actigraphy rest interval estimation, but can be impractical for large cohort studies and suffer from reproducibility problems. We developed a semi-automated pipeline (actiSleep) to set rest intervals consistent with best-practice hand-scoring algorithms incorporating event marker, diary, light, and activity data. To evaluate actiSleep performance, we used data from an observational study of 51 adolescents (14-19yr), with and without family history of bipolar disorder. Participants completed 2 weeks of wrist actigraphy and daily sleep diary. We first hand-scored records using a standardized hierarchical algorithm incorporating event marker, diary, light, and activity data. We then compared the hand-scored rest intervals to those from actiSleep and two automated activity-based algorithms ( Activity-Merged, Activity-Only). Activity-Only used activity-based sleep estimation and Activity-Merged joined closely adjacent rest intervals. For rest onset, rest offset, and rest duration, all algorithms had strong mean agreement with hand-scoring: actiSleep estimates were within 1-3 minutes, Activity-Merged within 2-4 minutes, and Activity-Only within 7-14 minutes. However, actiSleep had notably better (narrower) margins of agreement with hand-scoring, as evidenced by Bland-Altman plots, and greater positive predictive value and true positive rates for rest detection, especially in the 60 minutes surrounding the onset and offset of the rest interval. The actiSleep algorithm successfully estimates actigraphy rest intervals comparable to hand-scoring while avoiding pitfalls of activity-only algorithms. actiSleep has potential to replace hand-scoring for research in adolescents but requires further testing and validation in other samples.

18
Individual and Occupational Predictors of Sleep Intervention Response Among Shift-Working Firefighters

Song, Y. M.; Jeon, S.; Cho, A.; Chung, S.; Ahn, Y.; Suh, S.; Kim, J. K.

2025-10-29 psychiatry and clinical psychology 10.1101/2025.10.27.25338848 medRxiv
Top 0.1%
19.1%
Show abstract

BackgroundCognitive Behavioral Therapy for Insomnia (CBT-I) is the first-line treatment for chronic insomnia, yet its effectiveness varies substantially across individuals. This variability is particularly pronounced in shift workers, who often experience irregular sleep-wake cycles and high levels of stress. ObjectiveWe examined heterogeneous CBT-I response patterns and their predictors in a high-risk shift-working firefighter cohort, with the goal of identifying modifiable clinical and occupational factors that can inform personalized intervention strategies. MethodsOver a five-week CBT-I intervention, self-reported sleep quality was longitudinally assessed in 50 participants (mean age = 34.8 {+/-} 7.7 years; 88% male). Growth mixture modeling identified latent response trajectories, and multinomial logistic regression examined predictors including sleep parameters, insomnia severity, alcohol use, and work-related characteristics. A validated mathematical alertness model was used to objectively characterize physiological response patterns. ResultsThree response patterns emerged: steady-poor (24%), steady-moderate (50%), and improving (26%). Poor response was associated with emergency duty, longer sleep onset latency, and higher insomnia severity. Improved response was predicted by 3-day shift cycles, lower alcohol intake, and shorter shift-work duration. Mathematical model-predicted alertness profiles paralleled subjective response trajectories. ConclusionsCBT-I response among shift workers is highly heterogeneous and shaped by both clinical and occupational factors. Identifying these moderators can guide tailored behavioral interventions to improve treatment outcomes for individuals with insomnia in demanding work environments.

19
Updated U.S. census benchmark sleep dataset v1.1

Jones, A. M.; Sheth, B. R.

2025-12-29 health informatics 10.64898/2025.12.27.25343087 medRxiv
Top 0.1%
19.1%
Show abstract

We previously documented and released a benchmark dataset for machine learning research on sleep stage classification [1]. Subsequently, it was pointed out in a preprint [2] that some recordings in the National Sleep Research Resource [3] include only binary wake-sleep annotations, instead of full sleep stage scoring using the Rechtschaffen and Kales (R&K) [4] or American Academy of Sleep Medicine (AASM) [5] standards. Because wake-sleep labels are an ontological mismatch and not just label noise, they do not belong in a dataset designed for full sleep stage classification. Therefore, we have updated our benchmark dataset (henceforth known as benchmark v1.0) to replace 16 recordings with suitable recordings from age- and sex-matched subjects, while all other dataset selection criteria and distributions have been preserved. Additionally, the total number of recordings and the composition of the training, validation, and testing sets remain unchanged. While this update is a minor revision, we want to distinguish its use from v1.0, and therefore have titled this update as benchmark v1.1. The file listings are provided on the GitHub repository (https://github.com/adammj/ecg-sleep-staging/).

20
Performance of an Electroencephalography-Measuring Headband or Actigraphy Compared with Polysomnography in Older Adults with Sleep Disturbances

Miner, B.; Pan, Y.; Cho, G.; Talarczyk, J.; Chen, A.; Burzynski, C.; Polisetty, L.; Doyle, M.; Iannone, L.; Mejnartowicz, S.; Breier, R.; Gill, T. M.; Yaggi, H. K.; Knauert, M.

2025-01-27 geriatric medicine 10.1101/2025.01.25.25321124 medRxiv
Top 0.1%
18.5%
Show abstract

Study ObjectivesIn older adults, self-reported sleep measures may be inaccurate, but polysomnography (PSG) is burdensome. We assessed the performance of an electroencephalography-measuring headband (HB) or actigraphy (ACT) compared with PSG in older adults with sleep disturbances. MethodsSixty-three adults aged [&ge;]60 years who reported symptoms of insomnia and/or daytime sleepiness [&ge;]once/week completed a week-long, home-based protocol during which they wore the HB for seven nights, an actigraph for seven days and nights, and completed a one-night level II unattended PSG. For the current analysis, we compared total sleep time (TST) and wake after sleep onset (WASO) from all three devices on the PSG night. We calculated absolute differences and intraclass correlation coefficients (ICCs) for TST and WASO between HB and ACT, respectively, vs. PSG. We also evaluated the performance of the HB among subgroups of the poorest sleepers according to the presence of sleep apnea, insomnia, poor sleep quality, and periodic limb movements of sleep. Feasibility of the HB was assessed by measures of adherence (i.e., ability to use the HB over seven nights) and usability (i.e., ratings of items from the WEarable Acceptability Range [WEAR] scale). ResultsThe average age was 72.8 [standard deviation 6.6] years, 63.5% were female, and 63.5% identified as non-Hispanic White. On PSG, averages for TST and WASO were 370.1 [93] and 88.9 [63] minutes, respectively. For the HB vs. PSG, mean differences and ICCs were -11.9 minutes and 0.83 [0.74, 0.89] for TST; and -15.5 minutes and 0.65 [0.48, 0.77] for WASO. For ACT vs. PSG, mean differences for TST and WASO were larger, and ICCs showed lower levels of agreement. The HB performed well among the poorest sleepers, with ICCs >0.65 for TST and WASO. On average, participants wore the HB for 6.5 [0.8] nights, and usability was rated highly. ConclusionsThe HB demonstrated good agreement with PSG, outperforming ACT, including among the poorest sleepers. Devices like the HB might provide feasible measures of sleep that are more accurate than ACT and enhance the management of sleep health in older adults with sleep disturbances. Future research should focus on further validation of these devices in habitual sleep environments.